Papers with assessment methodology
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems (2024.acl-long)
Copied to clipboard
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, Maosong Sun
| Challenge: | Large Language Models (LLMs) and Large Multimodal Models have exceeded general human capabilities in various tasks. |
| Approach: | They present an Olympiad-level bilingual multimodal scientific benchmark featuring 8,476 problems from Olympiad level mathematics and physics competitions. |
| Outcome: | The best performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning. |
Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites (D19-1)
Copied to clipboard
| Challenge: | Currently, standard methods for style transfer have several significant problems. |
| Approach: | They propose to take BLEU between input and human-written reformulations into consideration for benchmarks. |
| Outcome: | The proposed architectures outperform state-of-the-art in style transfer metric on human-written reformulations and take BLEU between input and output into consideration for benchmarks. |